Back

American Journal of Epidemiology

Oxford University Press (OUP)

Preprints posted in the last 90 days, ranked by how well they match American Journal of Epidemiology's content profile, based on 67 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit.

1
its2s: a Python package for two-stage interrupted time series analysis using machine learning

Wilner, L.; Casey, J. A.; Mooney, S. J.; Do, V.; Ma, Y.; Benmarhnia, T.; Dey, A. K.

2026-07-06 epidemiology 10.64898/2026.07.02.26357175 medRxiv
Top 0.1%
34.6%
Show abstract

When randomized controlled trials are infeasible, researchers may leverage natural experiments for causal inference. Interrupted time-series (ITS) designs compare observed post-event trends to counterfactual predictions from pre-event data. Two-stage ITS designs use flexible models to generate optimized counterfactual predictions in the first stage, then estimate intervention effects by comparing observed to predicted outcomes in the second stage. Fitting high-dimensional versions of these models is challenging, requiring systematic infrastructure to ensure rigor and reproducibility. In response, we developed its2s, an open-source Python package implementing the two-stage ITS design with machine learning. its2s allows users to specify an intervention date and training/testing periods, select among built-in model architectures (e.g., Prophet-XGBoost, NeuralProphet), and generate confidence intervals via moving block bootstrap, preserving temporal autocorrelation in residuals. its2s layers defaults, configuration files, and runtime overrides to support workflows ranging from rapid default implementations to highly tailored analyses. We validated its2s using two case studies: a simulation with a 12% policy effect, recovering the true effect as 11.77%, and an analysis of the 2021 Pacific Northwest heat dome, finding 53% excess injury mortality over the following three weeks. its2s provides a flexible, reproducible framework for ITS-based quasi-experimental research, lowering barriers to rigorous machine learning-based counterfactual modeling.

2
Bias from small-count suppression in county-level cancer disparity estimates: a calibrated simulation study

gahan, k.

2026-06-08 epidemiology 10.64898/2026.06.05.26355021 medRxiv
Top 0.1%
33.8%
Show abstract

Abstract Background. Area-level cancer disparities are routinely estimated from public county data in which rates based on small counts (fewer than 16 cases or deaths) are suppressed. Analysts typically drop suppressed counties (complete-case analysis). Because suppression depends on case counts tied to population size and demographic composition, this missingness may be informative, but its effect on the disparity estimate has not, to our knowledge, been quantified. Methods. In a cross-sectional ecological study of 3,143 U.S. counties (analytic sample 3,018 with computable exposure) using one frozen public release of NCI State Cancer Profiles incidence and mortality data and ACS 2018-2022 5-year data, we estimated the most- versus least-deprived ICE(race+income) quintile rate ratio (RR) and rate difference for female breast, stomach, and cervix cancers under four suppression-handling methods: complete-case, available-case, bounding, and model-based small-area estimation. We characterized which counties were erased, and, following the ADEMP framework, ran a Monte Carlo simulation (1,000 replicates per cell; Monte Carlo standard error of bias approximately 0.0025) calibrated to the release to measure bias against a known truth. Analyses were pre-registered. Results. The suppressed fraction rose with rarity: 7.4% of counties for breast, 61.3% for stomach, and 75.7% for cervix incidence. Suppression was concentrated in the most-deprived quintile (cervix, 81.8% suppressed vs 63.8% least-deprived) and overwhelmingly removed rural rather than minority residents (cervix: 81% of the rural but 9% of the minority population erased). For breast (little suppression) the RR was 0.87 (95% CI 0.85-0.89) and identical across methods; for cervix incidence the complete-case RR (1.56) exceeded the model-based estimate (1.50), and for cervix mortality (91% suppressed) complete-case (1.86) exceeded model-based (1.56) by 16% with a wide bounding interval (1.88-2.62). In calibrated simulation, population-weighted complete-case bias was small (less than 2%) at the observed deprivation-county-size correlation and grew with rarity, threshold, and unweighted aggregation; its direction was conditional, becoming positive (over-estimation) as deprived counties became smaller. Conclusions. Complete-case handling of suppressed counties over-estimates rare-cancer area disparities relative to methods that retain them, while silently erasing most of the rural and most-deprived communities the estimate is meant to represent. The effect is negligible for common cancers and grows with rarity. Public-data disparity analyses should report the suppressed fraction and use bounded or model-based estimates by default. Keywords: cancer disparities; small-count suppression; Index of Concentration at the Extremes; informative missingness; small-area estimation; rural health.

3
A Simulation Study Comparing Multiple Imputation and Complete Case Analysis for Handling Missing Preschool Body Mass Index

Savu, A.; Dover, D. C.; Hajihosseini, M.; Gaudet, L. A.; Kaul, P.

2026-08-14 epidemiology 10.64898/2026.08.13.26360115 medRxiv
Top 0.1%
31.3%
Show abstract

Background and Objective. Missing data frequently occurs in health databases and can bias analyses if not correctly dealt with. Using real-world data, we compared complete-case and multiple-imputation methods for recovering true parameters of a multivariable logistic regression model for the association between maternal glucose levels during pregnancy and child excess weight at preschool age, where missing values were present in as much as 30% of our sample. Methods. This study utilized a cohort of 130,424 children with complete preschool-age body mass index (BMI) measurements from the Calgary and Edmonton health regions of Alberta, Canada. In the complete BMI data, we introduced missingness through deletion following three distinct mechanisms: missing completely at random (MCAR), at random (MAR), and not at random (MNAR). To handle the missing data created, we employed complete-case and multiple-imputation methods. Maternal glucose levels during pregnancy were categorized into five groups and its association with child excess weight at pre-school age was determined based on a logistic regression model using the full observed data (yielding true values), observed data that was not deleted (complete-case estimates), and imputed data (multiple-imputation estimates). The accuracy of complete-case and multiple-imputation estimates were evaluated against the true values. Finally, we conducted a sensitivity analysis for the MNAR mechanism using pattern-mixture models with an additive shift. Results. Under MCAR and MAR, multiple-imputation generally outperformed complete-case, yielding smaller absolute and relative bias. Both methods achieved high significance ([≥] 0.96) for most effects. Mean squared errors for multiple-imputation and complete-case were similar missing completely at random, missing at random, and coverage was consistently high ([≥] 0.99). Under MNAR, both complete-case and multiple-imputation showed poor performance regarding bias and statistical significance. Sensitivity analysis using pattern-mixture models indicated performance varied by specific effect. Conclusions. Under MCAR and MAR, multiple-imputation introduced higher bias but demonstrated superior overall performance based on mean squared error and restored statistical power. Conversely, both methods failed under MNAR, where pattern-mixture modeling sensitivity analyses revealed highly variable, effect-specific performance due to unverifiable shift assumptions. When faced with missing data, researchers should assess missingness mechanisms, report both complete-case and multiple-imputation estimates under MCAR/MAR while accounting for power-versus-bias tradeoffs, and employ pattern-mixture sensitivity analyses to test robustness when MNAR is plausible.

4
US County-level Structural Racism Effect Index and Cardiovascular Disease Mortality among Older Adults: A Bayesian Spatiotemporal Modeling

Begum, T.; Shahjahan, M.; Chakraborty, H.

2026-07-13 epidemiology 10.64898/2026.07.10.26357792 medRxiv
Top 0.1%
23.0%
Show abstract

Background: Cardiovascular disease (CVD) remains the leading cause of mortality among older U.S. adults, yet the contribution of neighborhood-based structural racism remains inadequately quantified. This study quantifies the association between the Structural Racism Effect Index (SREI) and CVD mortality among adults aged {greater than or equal to}65 years, evaluating how this relationship varies across U.S. geographic regions to identify key areas for intervention. Methods: This ecological study applied a hierarchical Bayesian spatiotemporal framework to 2017-2020 Centers for Disease Control and Prevention (CDC) Wide-Ranging Online Data for Epidemiologic Research (WONDER) data to estimate the association between SREI and CVD mortality across 3,007 U.S. counties. SREI was modeled continuously and categorically, adjusting for sociodemographic covariates. Population attributable fractions (PAF) and attributable deaths (AD) quantified the potentially preventable burden and its spatial disparities. Results: From 2017 to 2020, approximately 2.79 million CVD deaths were observed, with significant spatial clustering (Moran's I = 0.35, p < 0.001). Each standard-deviation increase in SREI was associated with 13% higher CVD mortality (IRR: 1.13, 95% CrI: 1.12-1.15). A positive dose-response gradient was observed across SREI quartiles, with mortality 24% higher in the highest quartile than in the lowest (IRR: 1.24, 95% CrI: 1.20-1.28). The PAF was 6.94% (95% CrI: 6.13-7.73), corresponding to 193,472 potentially preventable deaths. High exceedance probabilities (>0.95) were concentrated in the Southeast, Appalachia, and the Midwest. Conclusions: Structural racism is a spatially patterned, dose-dependent predictor of older adult CVD mortality, underscoring the need for public health monitoring and neighborhood-based upstream interventions where disease burden is concentrated. Keywords: Structural Racism Effect Index; Neighborhood disadvantage; Cardiovascular Disease Mortality; Bayesian Spatiotemporal Analysis; Population Attributable Fraction; Health Disparities; Health Equity.

5
Describing health inequalities without distortion: Simple-Means MAIHDA vs Random-Effects MAIHDA

Merlo, J.; Bashir, N. Z.; Rodriguez-Lopez, M.; Khalaf, K.; Öberg, J.; Perez-Vicente, R.

2026-08-18 epidemiology 10.64898/2026.08.17.26360592 medRxiv
Top 0.1%
22.1%
Show abstract

Multilevel Analysis of Individual Heterogeneity and Discriminatory Accuracy (MAIHDA) describes health inequalities through three components: (i) specific contextual effects (SCE), (ii) general contextual effects (GCE), and (iii) discriminatory accuracy of the context. We present Simple-Means MAIHDA (S-MAIHDA), which estimates each stratum directly from its observed individuals, with no distributional assumption. The observed proportions are unbiased whatever the stratum size, and their confidence intervals report the uncertainty honestly. S-MAIHDA operationalises the three components on the probability scale. The SCE are the raw and standardised stratum prevalences and the modification of the sociodemographic average differences by the area. The GCE are the variance partition coefficient (VPC) and the contextual structuring of the between-stratum inequality, expressed as the contextual clustering of inequalities, the additive sociodemographic differences, and the contextual modification of inequalities (CMI). The contextual discriminatory accuracy is expressed by the area under the ROC curve (AUC), and the sensitivity and specificity at the population prevalence as the threshold for a possible intervention. Because its estimates are the observed data themselves, S-MAIHDA is the canonical description, and the compare diagnostic quantifies how Random-Effects MAIHDA (RE-MAIHDA), the usual implementation, departs from it: RE shrinkage pulls small strata towards the overall mean and can hide the very inequalities the analysis seeks. The approach is implemented in the smaihda Stata command and reproduced in free Python code. We illustrate S-MAIHDA on register data from Malmo, Sweden (43,291 individuals; 300 area-sociodemographic strata), showing how the three components separate two contrasting outcomes: psychotropic medication use, almost purely sociodemographic, stable across areas, with weak contextual structuring (VPC {approx} 4%, CMI {approx} 0%); and choice of a private general practitioner, strongly geographical (VPC {approx} 11%, CMI {approx} 17%), with the sociodemographic differences reshaped and amplified in wealthy areas. RE-MAIHDA attenuated inequalities. For describing inequalities, S-MAIHDA preserves what the data show.

6
Predicting Subjective Cognitive Decline on Future BRFSS Survey Years: An Open Multi-Language Machine Learning Benchmark

Nguyen, T. T.; Nguyen, T. D.

2026-08-21 public and global health 10.64898/2026.08.18.26360709 medRxiv
Top 0.1%
19.5%
Show abstract

Background and Objectives: Subjective cognitive decline (SCD), self-reported worsening confusion or memory over the past year, is a common early marker of cognitive concern with relevance for Alzheimer's disease prevention and population health. Population-based machine learning benchmarks that respect temporal drift in public health surveillance remain limited. We developed a reusable multi-language prediction and interpretability framework for SCD using Behavioral Risk Factor Surveillance System (BRFSS) Cognitive Decline data. Methods: We analyzed pooled national (n = 298,944) and New York (n = 30,366) cohorts with chronological train (2015-2019), validation (national: 2020-2022; New York: 2020-2021), and locked test (2023-2024) splits. Nested LASSO identified stable predictors. Sixteen machine learning algorithms were compared under year-grouped cross-validation with SMOTE restricted to training folds. Four end-to-end Python/R pipelines (single-model or soft-voting) used validation-only isotonic calibration and Youden thresholding. Primary reporting pipelines were prespecified before test unlock (national: R tidymodels single-model; New York: Python single-model); algorithms within each pipeline were chosen by validation ROC-AUC. Post-hoc GLMs (national unweighted; New York design-weighted) and two training-only knowledge-graph layers supported interpretability. Results: Locked-test discrimination was consistent across implementations (ROC-AUC approximately 0.76-0.77). Prespecified pipelines achieved test ROC-AUC 0.770 (95% CI 0.767-0.773) nationally (R gradient boosting) and 0.762 (95% CI 0.746-0.777) in New York (Python AdaBoost). Soft-voting pipelines performed similarly (national 0.770; New York 0.757) and were treated as sensitivity benchmarks. Predicted probabilities were reasonably calibrated (Brier 0.118 nationally; 0.112 in New York), and higher scores among SCD-positive respondents persisted across survey years. Difficulty deciding, mental health, and functional health items ranked highest across permutation importance, SHAP, and GLMs. Respondents who reported no difficulty deciding (DECIDE = 2) had substantially lower odds of SCD than those who reported difficulty (DECIDE = 1; aOR approximately 0.13; FDR < 0.05). Training-only knowledge graphs likewise placed difficulty deciding nearest to SCD in both cohorts. Conclusions: A temporally locked, multi-pipeline BRFSS benchmark yields stable future-year SCD risk ranking, usable calibrated probability scores that remain separated by SCD status across survey years, and convergent interpretability signals. The open implementation supports reproducible surveillance-oriented machine learning for cognitive health.

7
Death in People with Down syndrome: Mortality statistics and novel predictors in US Medicaid and Medicare enrolled adults.

Tewolde, S.; Rosellini, A. J.; Michals, A.; Skotko, B. G.; Fortea, J.; Khor, B.; Handelman, S.; Rubenstein, E.

2026-07-20 epidemiology 10.64898/2026.07.17.26358090 medRxiv
Top 0.1%
19.5%
Show abstract

People with Down syndrome have higher age-specific mortality rates compared to the general population as well as peers with other intellectual and developmental disabilities. While a large proportion of mortality is attributable to Alzheimers disease, many die prior to Alzheimers diagnosis and some live to old ages, dying without Alzheimers. Our objectives were to use 11 years of Medicaid and Medicare data to describe characteristics and factors related to death in adults with Down syndrome and use machine learning to identify which conditions most strongly predict death in the full population and stratified by age. We identified death using Center for Medicare and Medicaid Systems reported date of death health conditions using ICD 9 and 10 codes. We used a case-control design with risk set sampling to have that controls to mimic the distribution of times of incident Alzheimers disease. We trained gradient boosted trees to identify strongest predictors. Our cohort included 137,293 adults with Down syndrome. Among those, 30,894 (22.5%) died during the study period. Mean age at death among those who died was 55 years (SD=10). Mean age of death in those with Alzheimers disease was 59 (SD=7) and those without was 52 (SD=12). The most influential predictors of mortality were any claim for dementia, any claim for pneumonia, re-occurring claim for cardiovascular disease three years before index death, and any claim for heart failure and epilepsy. Our results align with previous clinical work and highlight intervenable areas to reduce mortality in the Down syndrome population.

8
Simulation of synthetic health records for assessment of causal inference methods for vaccine efficacy

Velasco Pardo, V.; Daines, L.; Katikireddi, S. V.; Ritchie, L.; Robertson, C.; Simpson, C. R.; McCowan, C.; Swallow, B.

2026-07-19 infectious diseases 10.64898/2026.07.17.26358308 medRxiv
Top 0.1%
19.4%
Show abstract

Background During the COVID-19 pandemic, public health agencies used near real-time observational data to answer questions regarding vaccine effectiveness. However, traditional observational methods do not allow conclusions regarding counterfactual scenarios to be drawn from clinical data. Counterfactuals, which are outcomes that would have occurred under alternative interventions, can be used to formally assess the causal effects of public health interventions on health outcomes while accounting for the effects of confounding. Ideally individual patient data is used for the development of counterfactuals. Low-fidelity synthetic data may be useful for advancing methodological development where governance and privacy constraints prohibit access to sensitive personal data. Methods We simulated synthetic datasets based on the EAVE-II COVID-19 platform which has been limited to use for surveillance purposes. EAVE-II includes almost all resident people in Scotland registered with qualified general medical practitioners. Patient characteristics were simulated to reflect the known distribution of the Scottish population, accounting for dependencies between variables. Each synthetic dataset was encoded to different realistic scenarios for EAVEII 'ground truth' vaccine rollout and effectiveness results, explicitly stating the causal and confounding mechanisms, using a statistically sound method based on marginal structural models. Synthetic datasets of 100,000 individuals were then generated across five confounding scenarios and five severe outcome types. Results In scenarios with weak confounding, both unweighted and inverse probability of treatment weighted (IPTW) logistic regression recovered the true causal parameters. As confounding strength increased, only weighted models recovered the true mechanism. Conclusions Low-fidelity synthetic datasets simulated from EAVE-II data analysts to build and test causal inference pipelines, develop novel analysis pipelines, and train new researchers while awaiting access to real data. We showed how to generate synthetic datasets from a marginal structural model under different confounding scenarios.

9
NeMMo: an improved statistical algorithm for excess all-cause mortality surveillance and monitoring

Lytras, T.; Athanasiadou, M.

2026-08-17 epidemiology 10.64898/2026.08.14.26360477 medRxiv
Top 0.1%
18.8%
Show abstract

Background: Reliable estimation of excess mortality is central to population health surveillance. We introduce NeMMo (New Mortality Model), an evolution of the EuroMOMO model for estimating weekly all-cause expected mortality, and assess its behaviour and performance on empirical data. Methods: NeMMo incorporates population offsets, stratifies observed deaths by age group and models seasonality using a periodic B-spline rather than a Serfling-type sinusoidal function. Baseline weeks are selected by a data-driven procedure minimizing the skewness of the residuals before refitting the model, instead of relying solely on fixed calendar windows. NeMMo enables pooling across age groups, direct age standardization and incorporation of external predictors. We applied NeMMo and EuroMOMO to mortality and population data downloaded from Eurostat for 31 countries from 2015 onwards, excluding the COVID-19 pandemic period from baseline estimation. Results: For most countries NeMMo produced a higher expected mortality baseline that better tracked observed deaths, as well as tighter prediction intervals and higher maximum Z-scores, suggesting improved discrimination of mortality excesses. Z-scores and P-scores during non-pandemic weeks were closer to zero with NeMMo than with EuroMOMO but further elevated during pandemic weeks, providing greater separation between pandemic and non-pandemic mortality. Incorporating population offsets resulted in negative linear trends across all countries, consistent with declining mortality after accounting for demographic changes. The periodic B-spline identified substantial heterogeneity in the shape and timing of seasonal mortality that was not captured by a sinusoidal function. Conclusions: NeMMo provides a flexible and parsimonious framework for all-cause mortality surveillance that improves the established EuroMOMO model and offers theoretical, empirical and practical advantages. It is thus suitable both for detecting short-term spikes and for the long-term, age-adjusted quantification and comparison of mortality excesses that has become increasingly important since the COVID-19 pandemic. The accompanying 'nemmo' package for R facilitates its widespread adoption and application.

10
Minimal measurement strategies for cardiometabolic risk classification within the Positive Health framework: A cross-sectional analysis of NHANES 1999-2004 data

Schorr, K.; van den Broek, T.; van den Eijnden, M.; Hoevenaars, F.; Wopereis, S.

2026-08-10 epidemiology 10.64898/2026.08.07.26359959 medRxiv
Top 0.1%
18.8%
Show abstract

Background: Large-scale prevention and population health monitoring require measurement approaches that are both feasible and informative. Although several self-measurable anthropometric and fitness indicators have been associated with cardiometabolic risk, it remains unclear whether combining multiple measurements provides meaningful improvements over simpler approaches. We evaluated whether a parsimonious set of self-measurable indicators can achieve classification performance comparable to a full candidate set and quantified the incremental value of additional measurements. Methods: Using data from 8,275 adults in the NHANES 1999-2004 cohorts, we evaluated a predefined minimal set of four self-measurable anthropometric and fitness indicators (body mass index (BMI), waist-to-height ratio (WHtR), mid-upper arm circumference (MUAC), and VO2max (as a proxy for the 6-minute walk test) as candidate indicators of cardiometabolic risk. Their ability to reflect underlying clinical risk factors related to adiposity, glucose and lipid metabolism, and physical fitness was assessed using nested logistic regression models, likelihood ratio tests, discrimination metrics, and decision tree analyses. Results: WHtR consistently showed the strongest discriminative performance, with {Delta}PR-AUC values for BMI versus WHtR ranging from -0.002 to -0.037, and emerged as the primary splitting variable. Adding BMI to WHtR resulted in small gains in PR-AUC for most outcomes, ranging from 0.000 to 0.008, except for triglycerides where the gain was larger ({Delta}PR-AUC=0.039). Further inclusion of MUAC and VO2max provided limited additional value overall, with evidence of variation across outcomes and sex stratified analyses. Conclusion: Most classification performance was achieved using a limited number of simple self-measurable indicators, with little additional benefit from incorporating further measurements. These findings suggest that parsimonious measurement strategies may provide a feasible approach for cardiometabolic risk classification in population health and prevention settings while reducing measurement burden.

11
County Year Informatics Model for Annual and Cumulative Unique Lung Cancer Screening Eligibility in Maryland, 2026 to 2045

Adebamowo, C.; Adebamowo, S. N.

2026-06-17 epidemiology 10.64898/2026.06.15.26355716 medRxiv
Top 0.1%
15.5%
Show abstract

Purpose: Population-level lung cancer screening programs require denominators that reflect age, smoking history, geography, and changing eligibility over time. We estimated annual prevalent and 20-year cumulative unique low-dose computed tomography screening eligibility for Maryland residents under alternative screening criteria. Methods: We built a deterministic cohort-cell stock-flow simulation using Maryland county-equivalent jurisdiction projections by age, sex, and race/ethnicity, with ACS socioeconomic/nativity covariates and smoking-history priors for ever-smoked status, pack-years, and quit-years. Scenarios included USPSTF 2013 legacy, USPSTF 2021, ACS 2023/2024, a risk-model-expanded sensitivity, and ever-smoked-only capacity stress tests. Cumulative unique eligibility counted people once at first eligibility rather than summing annual prevalent person-years. Results: Under USPSTF 2021, an estimated 238,346 Maryland residents were eligible in 2026 and 245,326 in 2045. The 20-year cumulative unique denominator was 768,668, whereas naively summing annual prevalent counts produced 4,850,735 person-years, a 6.31-fold overcount. ACS 2023/2024 expanded annual eligibility to 314,616 in 2026 and cumulative unique eligibility to 902,796 by adding remote former smokers. Ever-smoked-only adult eligibility was 1,957,699 in 2026 and 3,383,683 cumulative unique over 20 years. Conclusion: A Maryland statewide screening initiative should plan from cumulative unique eligibility and county-equivalent jurisdiction-specific burden rather than annual prevalence alone. Explicit pack-year and quit-year modeling materially changes statewide and county allocation compared with current-smoking proxy models.

12
AI-Assisted Longitudinal Analyses of Environmental and Psychosocial Determinants of Subjective Cognitive Difficulties

Ma, S.; Cao, C.

2026-06-22 epidemiology 10.64898/2026.06.18.26355982 medRxiv
Top 0.1%
15.5%
Show abstract

Short-term environmental exposures have been linked to cognitive and behavioral outcomes, although many reported associations may reflect broader geographic and contextual differences. Using longitudinal data from the All of Us Research Program (2018--2024), we linked daily weather and air-pollution exposures to repeated attention-related and subjective cognitive outcomes. Associations were evaluated using pooled, fixed-effects, lagged, and event-study analyses. Additional machine-learning analyses were conducted to explore potential heterogeneity and latent psychosocial structure. Replication analyses were performed using the 2024 Behavioral Risk Factor Surveillance System (BRFSS). Several environmental exposure measures showed small associations with cognitive outcomes in pooled analyses, but most attenuated substantially after accounting for within-location temporal variation. Mediation, sensitivity, and machine-learning analyses yielded similar conclusions. In contrast, mental-health burden, loneliness, and social functioning were consistently associated with subjective cognitive difficulty and exhibited substantially larger effect sizes than environmental exposures. Similar patterns were observed in BRFSS. Exploratory AI-assisted analyses yielded findings broadly consistent with the primary longitudinal analyses. These findings suggest that short-term environmental perturbations may have limited associations with cognitive outcomes after accounting for within-location variation, whereas psychosocial factors appear to be more consistently associated with subjective cognitive burden.

13
Estimating age-specific heterogeneity in SARS-CoV-2 transmission from prospective longitudinal studies: the importance of correcting for study design

Chervet, S.; Layan, M.; Boëlle, P.-Y.; Guedj, J.; van der Werf, S.; Kerneis, S.; Sermet-Gaudelus, I.; Cauchemez, S.; Opatowski, L.

2026-08-10 epidemiology 10.64898/2026.08.06.26358866 medRxiv
Top 0.1%
15.2%
Show abstract

Longitudinal household studies, combined with mathematical modeling, are widely used to characterize the drivers of respiratory pathogen transmission, including the effects of age and symptoms. In practice, household recruitment protocols vary across studies, potentially introducing biases into observed data. However, these biases are typically overlooked in statistical inference, and their impact on parameter estimates remains unknown. Here, we use synthetic household outbreak data simulated under different recruitment protocols to evaluate how recruiting through infected children affects estimates of age-specific infectiousness and susceptibility. We show that, under child-based recruitment, the standard likelihood, which accounts only for transmission dynamics, leads to underestimating child infectiousness and overestimating child susceptibility by more than 30%. We then propose a novel estimation framework that explicitly incorporates the household recruitment process into the likelihood and show that it substantially reduces these biases. Applying this new approach to a French household study conducted during the COVID-19 pandemic, we estimated that children under 6 had 49% lower infectiousness than teenagers and adults during the Alpha wave, whereas no difference was observed during the Omicron wave. This study demonstrates that ignoring recruitment protocols can bias key epidemiological parameter estimates and highlights the importance of accounting for study design.

14
MASCOT-DS improves transmission dynamics inference by integrating multiple epidemiological data streams with phylodynamic inference

Weidemueller, P. H.; Esquivel Gomez, L. R.; Rodriguez-Barraquer, I.; Mueller, N. F.

2026-08-25 epidemiology 10.64898/2026.08.21.26361056 medRxiv
Top 0.1%
15.1%
Show abstract

Tracking how an infectious disease spreads in time and space relies on several distinct sources of surveillance data, reported case counts, viral concentrations in wastewater, seroprevalence surveys, and pathogen genomic sequences, each of which is imperfect and captures only part of the underlying transmission process. These data streams are typically analyzed separately or with highly parameterized, disease-specific models, making it difficult to combine their complementary strengths. Here we present MASCOT-DataStreams (MASCOT-DS), a BEAST2 software package that extends the structured coalescent model MASCOT to jointly infer prevalence over time and transmission rates between locations from any combination of case counts, wastewater concentrations, seroprevalence surveys, and pathogen phylogenies. Using simulated outbreaks in structured populations, we show that MASCOT-DS accurately recovers true prevalence trajectories and between-location migration rates. We then apply MASCOT-DS to genomic, case count, wastewater, and seroprevalence data from the SARS-CoV-2 Epsilon wave (winter 2020-21) in three San Francisco Bay Area counties, reconstructing county-level prevalence dynamics and quantifying transmission within and into the region. By systematically removing individual data streams, we find that genomic data are uniquely required to estimate transmission between locations, while seroprevalence data are essential for anchoring the overall magnitude of an outbreak; case counts and wastewater concentrations play largely interchangeable roles in capturing outbreak shape. These results demonstrate that integrating complementary epidemiological data streams substantially increases the certainty of transmission dynamics estimates compared to relying on any single data stream, and provides a framework for evaluating the added value of different surveillance strategies.

15
Psychosocial Factors Outweigh Short-Term Environmental Exposures in Subjective Cognitive Difficulties: A Causal AI Study

Cao, C.; Ma, S.

2026-06-25 epidemiology 10.64898/2026.06.23.26356240 medRxiv
Top 0.1%
13.6%
Show abstract

Short-term environmental exposures have been linked to cognitive and attention-related outcomes, but the robustness of these associations remains uncertain. We linked daily weather and air-pollution exposures to repeated measures of subjective cognitive difficulties and attention-related outcomes among participants in the All of Us Research Program from 2018 to 2024. Associations were evaluated using complementary longitudinal and causal-inference approaches, including fixed-effects, lagged-exposure, and event-study analyses. Machine-learning methods were used to characterize heterogeneity and latent psychosocial structure, and findings were independently evaluated using 2024 Behavioral Risk Factor Surveillance System data. Several environmental exposure measures were associated with cognitive outcomes in pooled analyses; however, most associations attenuated substantially after accounting for within-location temporal variation. In contrast, mental-health burden, loneliness, and impaired social functioning remained consistently associated with subjective cognitive difficulty across analytical approaches. Similar patterns were observed in the validation dataset. These findings suggest that some observed environmental associations may reflect broader geographic and contextual differences rather than short-term environmental effects. Overall, psychosocial factors demonstrated more consistent associations with subjective cognitive difficulties than short-term environmental exposures across multiple analytical frameworks and independent datasets.

16
Heavy metal exposure and conditional survival time in U.S. adults: a censored quantile regression cohort study

Fang, X.; Schwartz, J.

2026-07-09 epidemiology 10.64898/2026.06.29.26356268 medRxiv
Top 0.1%
13.6%
Show abstract

Abstract Background. Chronic low-level exposure to lead, cadmium, mercury, and arsenic remains a determinant of premature mortality in the U.S. general population, but previous hazard-ratio analyses do not characterize how exposure shifts the lower tail of the survival distribution, where premature mortality is concentrated. Objectives. We estimated the association of whole-blood lead, whole-blood total mercury, urinary cadmium, and the sum of urinary inorganic and methylated arsenic species with the 10th, 25th, and 50th conditional quantiles of follow-up time to all-cause mortality among U.S. adults aged 40 years and older. Methods. NHANES Continuous 1999 to 2018 was linked to the National Death Index through December 31, 2019 (n = 29,652). Censored quantile regression was fit per metal on the log2 scale at quantiles {tau}{0.10, 0.25, 0.50}. A restricted-cubic-spline (RCS) censored-quantile-regression was fit for blood lead and urinary cadmium to investigate the threshold effect. Results. Over a median follow-up of 9.1 years, 7,215 deaths were ascertained. A doubling of urinary cadmium was associated with -1.57 years of follow-up (95% CI: -2.08, -1.07) at the 10th conditional quantile, -1.50 (-2.04, -0.96) at the 25th, and -1.49 (-1.93, -1.04) at the median (Benjamini Hochberg q < 0.001 throughout). A doubling of whole-blood lead was associated with -0.70 years (95% CI: -0.99, -0.40) at the 10th conditional quantile, -0.62 (-0.92,-0.31) at the 25th, and -0.61 years (-0.89, -0.34) at the median; the absolute loss was largest at {tau} = 0.10 for both metals. Urinary arsenic-metabolite sum was not associated with conditional follow-up at the estimable quantiles. Despite adjustment for dark and fatty-fish intake or DHA/EPA, whole-blood total mercury was associated with longer follow-up (i.e., negatively associated with mortality risk), possibly due to residual confounding by broader dietary or socioeconomic factors, rather than a true protective effect. The cadmium association was additionally robust to the mutual adjustment of lead. Discussion. Low-to-moderate urinary cadmium and whole-blood lead were associated with fewer years of follow-up survival at the lower-tail and median conditional quantiles of survival, with the largest absolute losses at the lower tail of the conditional survival distribution, where premature mortality is concentrated. These findings support continued reductions in U.S. cadmium exposure and lead with particular benefit for adults most vulnerable to premature death.

17
Joint Heat and PM2.5 Exposure Across US Metropolitan Areas: Multi-Stressor Disparities, Historical Redlining, and a Multi-Metric Assessment Framework

Mandalapu, S. V.; Sharma, R.; Pillarisetti, A.

2026-08-23 epidemiology 10.64898/2026.08.20.26360970 medRxiv
Top 0.1%
13.4%
Show abstract

Many urban health outcomes are shaped by environmental stressors that occur together rather than in isolation, yet methods for measuring such co-occurrence at the neighbourhood scale remain underdeveloped. We developed a multi-metric framework for joint co-exposure assessment and applied it to characterise the joint spatial distribution of summer surface heat and fine particulate matter (PM2.5) across 42,304 census tracts in 48 large US metropolitan areas during summers 2015 to 2020, covering approximately 174.6 million residents. The framework combines a composite co-exposure index, a joint exceedance indicator, a conditional exceedance ratio that compares observed joint occurrence to within-group statistical independence, and an upper tail dependence parameter estimated using both the non-parametric Caperaa-Fougeres-Genest estimator and a Gumbel copula, with bias-corrected and accelerated (BCa) confidence intervals obtained from a 5,000-replicate metropolitan-area block bootstrap. Among residents of predominantly Black tracts, 13.21% lived in neighbourhoods that simultaneously exceeded the within-metropolitan-area 80th percentile for both heat and PM2.5, compared with 3.33% of residents of predominantly White tracts; the corresponding heat-only and PM2.5-only ratios were 2.88 and 2.48. Residents of Home Owners Loan Corporation grade D tracts had 3.97 times the odds (95% confidence interval 2.79 to 5.66) of joint hotspot residence compared with grade A residents after adjustment for contemporary tract racial composition, poverty, renter-occupancy, and pre-1960 housing. The within-group conditional exceedance ratio at the 80th percentile was 2.29 in predominantly White tracts (95% BCa CI 1.81 to 2.78), 1.27 in predominantly Black tracts (0.71 to 1.56), and 1.13 in predominantly Hispanic tracts (0.70 to 1.41); the White interval excluded one while the Black and Hispanic intervals included one, which we interpret as power-limited given fewer contributing CBSAs. Magnitudes attenuated under near-surface air temperature surfaces but the direction and statistical significance of the primary findings were preserved. The framework is portable to other compound-exposure questions and supports cumulative-impact assessment.

18
Multi-season evaluation and analysis of categorical trend forecasts of influenza hospital admissions in the United States

Davis, J. T.; Kaur, G.; Hines, A.; Ben-Nun, M.; Venkatramanan, S.; Brooks, L.; Mathis, S.; Ajelli, M.; Litvinova, M.; Kummer, A. G.; Ventura, P. C.; Mhade, S.; Weber, D.; Shemetov, D.; DeFries, N.; McDonald, D. J.; Yamana, T.; Zepeda-Tello, R.; Shaman, J.; Yaari, R.; Pei, S.; Webber, A.; Shandross, L.; Ray, E.; Wadsworth, S.; Niemi, J.; Redman, W. T.; Mullany, L.; Posner, R.; Mallela, A.; Lin, Y. T.; Hlavacek, W. S.; Smart, A.; Gill, A. A.; Drennan, A.; Fiebiger, B. J.; Miller, E. F.; Lee, J.; Mihaljevic, J. R.; Geist, K. A.; Baltz, M.; Bernik, O.; Truong, Y.-M. B.; Chen, Y.; Grosvenor, C. J.;

2026-09-02 epidemiology 10.64898/2026.08.31.26361843 medRxiv
Top 0.1%
13.2%
Show abstract

Forecasting influenza hospitalizations informs public health preparedness, yet questions remain about which types of forecasts best guide action. We evaluate categorical trend forecasts, which communicate probabilities of upcoming increases or decreases in epidemic trajectories, submitted to CDC's FluSight Forecasting Challenge between Fall-2024 and Spring-2026. Teams submitted probability distributions over five categories describing direction and magnitude of week-over-week changes in laboratory-confirmed influenza hospital admissions. We assessed performance using Ranked Probability Skill Score, Brier Skill Score, and measures of forecast-observation agreement. Most models outperformed an equal-probability baseline; the FluSight ensemble ranked among the top three in the 2024-25 and 2025-26 seasons. Forecasts were most accurate during stable periods and least during periods of rapid change, with most models underestimating observed trends. Conclusions were robust to choice of scoring metric and reference model. These results support categorical trend ensembles as an approach to communicating infectious disease forecasts that may inform public health decision-making.

19
Addressing Measurement Error of Machine-Learned Physical Activity in Nonlinear Dose-Response Survival Analysis: Development and Evaluation of Accelerated Failure Time, Spline, and Simulation-Extrapolation Method

Mamiya, H.; Zhang, Q.; Zhang, X.; Yan, Y.; Sharma, A.

2026-08-31 epidemiology 10.64898/2026.08.25.26361155 medRxiv
Top 0.1%
13.1%
Show abstract

Wearable (accelerometer) data and machine-learning allow objective assessment of the amount of daily physical activity. However, wearable-derived human activity is subject to measurement error. No studies have corrected the dose-response association between physical activity and survival time to chronic diseases, including cardiovascular disease (CVD). The objective is to estimate the measurement error-corrected association between CVD events and multiple measures of daily duration of light and total physical activity, derived from machine-learning and conventional accelerometer-processing methods. Our method combined an accelerated failure time model, spline, and simulation-extrapolation (SIMEX). The method recovered the true dose-response non-linear association in simulated data, while the naive model failed to capture it due to substantial attenuation. Application to the UK Biobank accelerometer cohort also showed an increased protective association of total physical activity after SIMEX correction (Time Ratio [TR] = 1.56, 95% CI: 1.28-1.82 vs. TR = 1.38, 95% CI: 1.24-1.54 for SIMEX-corrected vs. uncorrected dose-response association between the 95th and 5th percentiles of total activity), with a similar increase for light physical activity. Sensitivity analysis indicates that the female population experiences a substantially larger protective association after SIMEX correction than males. Dose-response survival analysis is a widely used analytical method in physical activity epidemiology and benefits from measurement error correction.

20
Surrogate Endpoint Evaluation with Causal Mediation: Lessons from the A4 Trial

Hoefen, E. J.; Flanders, M.; Gantenberg, J.; Hayes-Larson, E.; Crane, P. K.; Choi, S.-E.; Trittschuh, E. H.; Ackley, S.

2026-07-27 epidemiology 10.64898/2026.07.23.26358810 medRxiv
Top 0.1%
13.0%
Show abstract

Surrogate endpoints, or measures used in place of the true outcome of interest, have relevance across multiple disease areas. The Prentice Criteria, proposed in 1989, assess surrogacy by evaluating how the treatment's effect on the true outcome operates through the potential surrogate. Using the A4 Study of solanezumab, we evaluate multiple formulations of the Prentice Criteria using causal mediation. We estimated direct and indirect effects of solanezumab on cognitive decline through cerebral amyloid across different, but reasonable, measures of exposure, mediator, outcome, and covariate adjustment. Unsurprisingly given that solanezumab did not show benefit, estimated indirect effects were close to zero. These results provide little evidence of meaningful mediation for memory and global cognition, but precision varied substantially. Confidence interval widths varied by up to a factor of 17 across specifications. Causal mediation analysis of individual-level randomized trial data may contribute to quantitative surrogate endpoint evaluation, but our findings indicate this is only the case when analytic choices are biologically justified, prespecified, and interpreted with attention to uncertainty.